Papers with Linguistic Data Consortium

12 papers
Introducing NIEUW: Novel Incentives and Workflows for Eliciting Linguistic Data (L18-1)

Copied to clipboard

Challenge: a 2010 survey found that the language of the European Union, not even English, was not fully supplied . the absence of Language Resources stifles teaching and technology building, authors say .
Approach: They propose to harness the power of alternative incentives to elicit linguistic data and annotation . they also describe changes to the workflows necessary to collect data from workforces attracted by incentives .
Outcome: a new initiative to harness incentives to elicit linguistic data and annotation improves language resources . the NIEUW project is funded by the u.s. national science foundation .
Reflections on 30 Years of Language Resource Development and Sharing (2022.lrec-1)

Copied to clipboard

Challenge: Linguistic Data Consortium was founded in 1992 to solve the problem that limitations in access to shareable data was impeding progress in Human Language Technology research and development.
Approach: They review the roles of the Linguistic Data Consortium over the past 30 years after describing the conditions that lead to an HLT winter followed by a reawakening and an insatiable hunger for LRs.
Outcome: The authors review the roles of the Linguistic Data Consortium over the past 30 years and provide a preview into future plans.
CAMIO: A Corpus for OCR in Multiple Languages (2022.lrec-1)

Copied to clipboard

Challenge: CAMIO is a corpus of 70,000 images of machine printed text for optical character recognition (OCR) it covers 35 languages across 24 unique scripts.
Approach: CAMIO is a corpus of annotated multilingual images for optical character recognition . the corpus includes nearly 70,000 images of machine printed text .
Outcome: The corpus includes nearly 70,000 images of machine printed text . most images have been exhaustively annotated for text localization .
Laying the Groundwork for Knowledge Base Population: Nine Years of Linguistic Resources for TAC KBP (L18-1)

Copied to clipboard

Challenge: Knowledge Base Population (KBP) evaluations target information extraction technologies for knowledge bases comprised of entities, relations, and events.
Approach: They describe the linguistic resources provided by Linguistic Data Consortium for TAC KBP since 2009 . they highlight changes made to support evolving evaluation requirements .
Outcome: The evaluations have targeted information extraction technologies for the population of knowledge bases comprised of entities, relations, and events.
A 2nd Longitudinal Corpus for Children’s Writing with Enhanced Output for Specific Spelling Patterns (L18-1)

Copied to clipboard

Challenge: IQB study looks at reading, mathematics and spelling ability across different states.
Approach: They collect three longitudinal corpora of German school children's weekly writing in German and transcribe them into a corpus for research via Linguistic Data Consortium.
Outcome: The corpus of German school children's weekly writing in German was collected and transcribed.
Related Works in the Linguistic Data Consortium Catalog (2020.lrec-1)

Copied to clipboard

Challenge: Existing metadata standards for Related Works are used to define relations between language resources.
Approach: They describe the development and implementation of a Related Works schema and the steps to implementation.
Outcome: The proposed schema has been implemented in the Linguistic Data Consortium's (LDC) Catalog.
A Progress Report on Activities at the Linguistic Data Consortium Benefitting the LREC Community (2020.lrec-1)

Copied to clipboard

Challenge: Linguistic Data Consortium (LDC) activities include the collection, annotation, processing, distribution, archiving and curation of language resources.
Approach: a new report sketches the activities of a data center devoted to supporting the work of LREC attendees . 96 new corpora released in 2018-2020 to date, a technology evaluation campaign and innovations to advance methodology for language data collection and annotation.
Outcome: 96 new corpora released in 2018-2020 to date, new technology evaluation campaign and innovations to advance methodology of language data collection and annotation.
Morphological Segmentation for Low Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of annotated morphological data is described for the DARPA LORELEI Program . the data is annotating 9 low resource languages and root information for 7 of the languages .
Approach: This paper describes a new morphology resource created by Linguistic Data Consortium and the University of Pennsylvania for the DARPA LORELEI Program.
Outcome: The annotated corpus provides a gold standard for unsupervised morphological segmenters and analyzers . the language-specific annotation guidelines were language-independent, but included morphology paradigms and other specifications.
From ‘Solved Problems’ to New Challenges: A Report on LDC Activities (L18-1)

Copied to clipboard

Challenge: This paper reports on the activities of the Linguistic Data Consortium .
Approach: This paper reports on the activities of the Linguistic Data Consortium . it summarizes the over 100 Language Resources released since the last report .
Outcome: The report summarizes the over 100 Language Resources released since the last report . many of the LRs have been contributed by research groups around the world .
The SAFE-T Corpus: A New Resource for Simulated Public Safety Communications (2020.lrec-1)

Copied to clipboard

Challenge: Linguistic Data Consortium developed the SAFE-T Corpus to support the NIST OpenSAT evaluation series.
Approach: They introduce a new resource, the SAFE-T Corpus, designed to simulate first-responder communications by inducing high vocal effort and urgent speech with situational background noise.
Outcome: The SAFE-T Corpus was developed to support the NIST OpenSAT (Speech Analytic Technologies) evaluation series.
A Large Scale Speech Sentiment Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpus for sentiment analysis uses text inputs, but voice inputs are becoming more important as smart assistants and mobile voice control become more prevalent.
Approach: They propose to extend the Switchboard-1 Telephone Speech Corpus by adding sentiment labels from 3 different human annotators for every transcript segment.
Outcome: The proposed corpus contains 49500 labeled speech segments covering 140 hours of audio.
Spanless Event Annotation for Corpus-Wide Complex Event Understanding (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for annotating multilingual, multimedia data are limited by the availability of multilingual corpora for schema-based event representation.
Approach: They propose a new approach to event annotation to promote whole-corpus understanding of complex events in multilingual, multimedia data.
Outcome: The proposed method is part of the DARPA Knowledge-directed Artificial Intelligence Reasoning Over Schemas (KAIROS) Program.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations